Synthetic Biology
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match Synthetic Biology's content profile, based on 24 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Straub, G.; Aldrich, D.; Tobin, C.
Show abstract
The Modular Cloning (MoClo) and PhytoBrick standards have revolutionized plant synthetic biology by establishing a standardized, hierarchical assembly grammar. However, as the engineering of complex metabolic pathways, multi-trait stacks, and synthetic gene circuits expands, existing toolkits hit practical boundaries in assembly capacity and fixed grammars. To overcome these bottlenecks, we present MozClo, an expansion of the MoClo/PhytoBrick architecture. MozClo expands the standard Level 1 assembly framework to 10 positions using new L1 acceptors, end-linkers and dummy parts. We also identify and resolve a critical, sticky-end collision at L1 position 7 that has caused assembly failures during L2 cloning of large plasmids. To address commercial DNA synthesis length constraints and to lower cloning costs, we designed a universal 5-in-1 gene fragment multiplexing system. This architecture embeds up to five distinct parts flanked by orthogonal pairs of BpiI restriction sites into a single synthesized fragment, allowing them to sort independently into their respective L0 acceptor plasmids while maintaining complete modular flexibility of part types. Finally, we provide Level 2 cloning backbones with built in selection genes for common soybean transformation methods to facilitate downstream plant selection. Together, these advancements reduce DNA synthesis overhead and accelerate the construction of complex multigene payloads for plant biotechnology.
Tassinari, E.; Ives, L.; Hawkins, E.; Annese, D.; Fonseca, S.; Lan, Y.; Haerty, W.; Wojtowicz, E.; Grandellis, C.
Show abstract
High-quality plasmid DNA purification at high throughput remains a significant bottleneck in molecular biology and bioengineering. Current methods frequently fail to deliver sufficient yields of pure, transfection-grade DNA required for genetic engineering applications in mammalian cells. Here, we present a Biofoundry-based automated pipeline using the CyBio FeliX robotic liquid handling platform to rapidly purify plasmid DNA with minimal manual intervention. The protocol leverages Solid Phase Reversible Immobilisation (SPRI)-based magnetic bead technology to ensure consistency, scalability, and DNA purity suitable for downstream viral particle production and mammalian cell transfection. The pipeline supports flexible processing of between 8 and 96 samples per run, making it adaptable across a wide range of experimental scales. The protocol is openly available via Earlham Institute GitHub repository, enabling broad adoption across the bioscientific community and contributing to the growing toolkit of reproducible, scalable engineering biology workflows. In this work, we employed an integrated robotic pipeline to process 528 pooled DNA plasmids and built a Lentiviral DNA plasmid library for lineage tracing, validated the library by sequencing, and demonstrated efficacy in downstream mammalian cell transfection experiments.
Adamson, H. E.; McLellan, J. R.; Singhal, K.; Demirel, M. C.; Salis, H. M.
Show abstract
Genetic systems engineering is constrained by high DNA synthesis costs, assembly inefficiencies, and challenges in expressing complex proteins. To address these limitations, we developed a highly parallel, low-cost pipeline for the design, assembly, and functional screening of genetic systems, which we stress-tested on highly repetitive structural proteins, including spider silk, biocements, reflectins, and talins. The integrated pipeline combines computational genetic systems design, low-cost many-plasmid DNA assembly from oligopools, automated many-to-many mapping using nanopore sequencing data, and a label-free biosensor to measure single-cell protein expression levels. We applied this pipeline to build 240 plasmids, achieving an 88% success rate (up to 2000 bp) using standard clonal isolation and 58% assembly efficiency (up to 5600 bp) without selective DNA purification, while lowering material costs by up to 24-fold. We applied the biosensor to identify genetic factors that create distinct cellular subpopulations with varying protein expression levels. Overall, the integrated pipeline will dramatically lower the cost of high-throughput synthetic biology, while demonstrating how designing genetic systems to improve build efficiency ("design for build") and directly incorporating biosensors into genetic systems ("design for test") will greatly accelerate design-build-test workflows.
Vora, S.; Styczynski, M. P.
Show abstract
While in vivo synthesis of biologic therapeutics has been broadly successful, it is limited by biological constraints of the cells and by the complexity, time, and cost of implementing the pipeline from discovery through manufacturing. Cell-free expression systems (CFES), which use cellular transcription and translation machinery to express proteins in vitro, offer a promising alternative approach that could improve robustness and modularity in that pipeline. However, current benchmark CFES productivity is well below the theoretical capacity of the input nucleotides and amino acids. Efforts to address this issue are hindered by limited understanding of the extent of enzymatic activity in CFES beyond gene expression, as previous work has shown that metabolic enzymes in cell-free lysates cause substantial background metabolic activity that influences protein expression. Here, we hypothesized that the inflection point of protein expression is a critical timescale for CFES metabolism. We performed metabolomics characterization of CFES reactions, finding significant metabolic changes at the inflection point. Driven by these findings, we sought to identify supplements that could be added to the cell-free reaction to avoid metabolic limitations. We found that amino acid supplementation increased expression productivity and lifetime only when added after the inflection point, and actually hurt expression when added before the inflection point. We found similar supplementation timing impacts for some other metabolites as well. These findings show that endogenous metabolism and supplementation timing are deeply interconnected and are critical considerations in CFES optimization, and that metabolomics-informed fed-batch supplementation is a potentially valuable strategy to improve reaction productivity.
Cardenas Ramirez, P.; Smick, S.; Dey, S.; Niles, J. C.
Show abstract
Malaria is responsible for over half a million deaths each year. However, our understanding of malaria parasite biology is hampered by a lack of molecular tools, particularly at the level of transcriptional control. In light of this, we have created two orthogonal systems for inducible transcriptional repression in the malaria parasite Plasmodium falciparum using bacterial repressor proteins. We achieve 200- to 800-fold repression of expression, improving on previous attempts at transcriptional regulation by two orders of magnitude and outperforming gold standard translational/post-transcriptional regulation systems. We developed automated DNA design software to apply this tool to conditional regulation of native gene expression, validating essentiality and chemogenetic interactions with both two parasite lipid kinases and PfKelch13, which is associated with artemisinin resistance. These tools can advance our understanding and engineering of malaria functional genomics, drug mechanisms, and gene regulation.
Hasenklever, J. C.; Paderi, V.; Hasenklever, D.; Axmann, I. M.; Schipper, K.
Show abstract
BackgroundThe corn smut fungus Ustilago maydis is an important microbial model organism representing a genetically amenable and readily cultivable basidiomycete. Research in this fungus addresses a broad range of fundamental questions and its biotechnological exploitation is on the rise. Although genetic engineering in principle is well established, efficient methodology for synthetic biology approaches such as metabolic engineering or pathway transplantation has remained limited. ResultsHere, we present a comprehensive toolbox for U. maydis based on modular cloning and the characterization of more than 20 promoters. Careful comparative evaluation of insertion loci and terminator as well as reporter effects was conducted and a novel color-based strategy for straightforward genome integration was implemented. Moreover, the cloning and subsequent one-step integration of four transcriptional units into U. maydis was demonstrated by creating a "rainbow" strain producing four fluorescent proteins. ConclusionOverall, this next generation toolkit strongly advances genetic engineering and systems biology approaches in U. maydis, fostering its development into a valuable and competitive fungal chassis and prime model, particularly in applied research.
McLellan, J. R.; Salis, H. M.
Show abstract
T7 RNA polymerase is widely used to produce RNA using a canonical T7 promoter; however, it will also bind to low-affinity sites to generate cryptic transcription and produce RNA byproducts, which reduce full-length mRNA purity and yield. When manufacturing therapeutic RNAs for clinical applications, RNA byproducts must be removed using costly downstream purification and can cause adverse immunogenicity. To predict T7 transcription rates and reduce cryptic transcription, we designed 11588 T7 promoters and measured their mRNA levels, spanning a 6300-fold range within in vitro transcription reactions. We developed the T7 Promoter Calculator, a sequence-to-function machine learning model that predicts the T7 transcription rate on arbitrary DNA sequence across a 500-fold range with high accuracy (R2 = 0.80), accounting for both core and flanking motif sequences. We combined the model with generative design to remove low-affinity T7 sites from a therapeutic T7 expression system, resulting in a 2-fold increase in full-length mRNA purity. The automated design of T7 expression systems to remove undesired RNA byproducts increases mRNA purity and lowers downstream separation costs, while reducing adverse immunogenicity.
Dysart, M. J.; Fang, L.; Karinje, L. K.; Chappell, J.; Stadler, L. B.; Silberg, J. J.
Show abstract
TEXT ABSTRACTCatalytic-RNA (cat-RNA) expressed from mobile DNA can record cellular events, such as the uptake of plasmids via horizontal gene transfer, by splicing a barcode onto 16S ribosomal RNA (rRNA) - a system termed RNA addressable modification (RAM). However, scaling RAM to record multiple simultaneous biological events requires large numbers of orthogonal cat-RNA whose signals reflect the biological features under investigation rather than variability arising from the barcode sequence. Here, we explore how to design orthogonal cat-RNA to record information about multiple plasmid-encoded traits in parallel. We show that cat-RNA having tRNA-derived barcodes with sequence variation in the anticodon stem-loop present greater signal consistency within Escherichia coli than mRNA-derived barcodes. When orthogonal cat-RNA designs harboring tRNA-derived barcodes were evaluated in Vibrio natriegens and Pseudomonas putida, increased variance was observed compared with Escherichia coli. Nevertheless, the signal consistency was sufficient to use these orthogonal cat-RNAs to report on the relative activities of four promoters and two origins of replication by sequencing barcoded-rRNA derived from the three organisms. These results show how RAM can be multiplexed to report on mobile DNA features in microbial communities and illustrate the importance of accounting for variability in RNA outputs when designing and interpreting multiplexed RNA barcoding data. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=88 SRC="FIGDIR/small/738544v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@406ebaorg.highwire.dtl.DTLVardef@259751org.highwire.dtl.DTLVardef@1f1512corg.highwire.dtl.DTLVardef@8384b_HPS_FORMAT_FIGEXP M_FIG C_FIG
Gilmour, A. R.; Wei, Q.; Hellinger, J.; Kulhanek, D. L.; Jansen, Z.; Baumer, K. M.; Brodbelt, J. S.; Thyer, R.
Show abstract
Selenocysteine (Sec), the 21st amino acid, is a rare non-canonical amino acid that represents an attractive target for protein engineering due to its desirable chemical properties such as high affinity for metals, strong nucleophilicity, and reversible covalent bond formation. To bypass the natural constraints on Sec placement within proteins, several strategies have been developed to rewire the native translational machinery to enable site-specific incorporation. However, these usually abolish the quality control mechanism that excludes the serine-charged selenocysteinyl-tRNA (Ser-tRNASec), the immediate biosynthetic precursor, from translation resulting in heterogenous protein species. This challenge is confounded by a lack of genetic tools to accurately report the selenylation state of the tRNA pool as most are blind to competing process of Ser incorporation, which can only be observed using analytical methods. To resolve this issue, we have developed a new fluorescent reporter, Selenocysteine Adjusted Ratiometric Chromophore (SeARCh), which exhibits two distinct spectral outputs dependent on the incorporation of either Ser (red) or Sec (green). Using SeARCh, we define several factors which influence the observed Sec:Ser ratio and construct a new hybrid biosynthetic pathway with improved performance, achieving 90% Sec incorporation. Furthermore, SeARCh displays unusually complex mass spectra due to the isotope distribution of selenium and heterogenous nature of the protein in solution and we report specific methods to account for this behaviour and precisely quantify the rare Ser-containing species found at high Sec incorporation efficiencies. Our findings suggest that the equilibrium between selenoprotein and tRNASec expression levels is a key driver of incorporation efficiency and implies a process that is broadly biosynthetically constrained. Collectively these tools represent a significant advance in the metrology of selenocysteine biosynthesis and incorporation and can be used to inform and standardize future engineering efforts.
Peterson, R.; Parvez, S.; Marsh, M. C.; Owen, S. C.
Show abstract
The dCas9 system has rapidly been developed into many tools to explore different aspects of the human genome; the high binding specificity, coupled with the inactive nuclease enzyme, allows for precise recruitment of molecules to specific sequences of DNA. We sought to exploit these capabilities to create a tool for assessing the real-time proximity of two DNA sequences within a biological system. By incorporating aptamers into gRNAs, dCas9 molecules can be used to recruit the {beta}9 or {beta}10 strands of split-NanoLuc(R) to specific DNA sequences and quantify the proximity of those sequences based on their ability to complex with the luciferase fragment ({Delta}11S) and produce luminescence. While many tools exist to detect a single DNA sequence, this system is uniquely capable of assessing how two DNA sequences interact with each other. As expected, we found that the interaction of two dCas9 molecules was affected by their linear distance from each other on dsDNA. Surprisingly, we found that their interaction was also strongly influenced by rotational orientation, even for sequences that are close together in linear space. This finding indicates that dCas9 rotational alignment is an important consideration for designing dCas9 systems that target multiple DNA sequences simultaneously. Beyond the findings presented herein, we believe this DNA proximity detection tool has the potential to be adapted for applications involving the proximity and orientation of two DNA sequences.
Bull, T. A.; Farina, L.; Sutton, S.; Van Blair, J.; Carlsen, L.; Khakhar, A.
Show abstract
Precise temporal control of gene expression is critical for studying plant biology and engineering complex crop traits. Current systems enable chemically inducible regulation but rely on costly or agriculturally impractical inducers and lack the flexibility needed to regulate combinations of native loci and transgenes. In this work, we elucidate the design rules for control systems, based on Cas9 and Cre recombinase fused to the ecdysone receptor (EcR), which respond to a widely used agrochemical methoxyfenozide (MF). First, we validated the function of both circuits in transient assays and explored how transduction properties can be modulated by engineering nuclear trafficking dynamics. We next characterized both the Cas9-based and recombinase-based systems by using them to regulate fluorescent reporters in transgenic Arabidopsis thaliana plants. Here, we demonstrate systemic activation following root application of MF, validating the use of an agriculturally compatible inducer for whole-plant gene regulation. Finally, we validated the utility of the inducible Cas-based SynTF system to regulate multigene pathways and control both metabolic flux and developmental circuits. Together, these results establish design principles for agrochemical-inducible control systems and demonstrate their utility for temporally regulating plant phenotypes. These synthetic circuits provide a versatile framework for engineering complex traits using an agriculturally compatible inducer.
Cai, Z.; Sang, Y.; Xu, L.; Chang, Y.; Wong, N. M.; Zhu, J.; Chen, C.-T.; Bao, Z.
Show abstract
Tandem Interspaced Guide RNA (TIGR)-TIGR-associated (Tas) systems are a newly discovered family of ultracompact, modular RNA-guided DNA-targeting proteins that function without a protospacer adjacent motif (PAM) requirement. Their utility as genome engineering tools in microbes remains unexplored. Here, we report the first functional implementation of TIGR-Tas in Saccharomyces cerevisiae for genome engineering. We show that TasR from Parcubacteria (ParTasR) can be programmed by user-defined tigRNAs to generate targeted DNA double-strand breaks at yeast endogenous loci. By co-delivering ParTasR with customized tigRNAs and donor templates, we achieved precise gene fragment deletion and targeted codon substitutions at multiple genomic loci. The multiplex genome engineering capability of this TIGR-Tas system was demonstrated through high-efficiency multiplex gene disruption and chromosomal assembly of a lycopene biosynthesis pathway while inactivating an endogenous gene. This work establishes TIGR-Tas as a valuable addition to the yeast genome engineering toolbox, particularly for applications requiring PAM-independent targeting or compact delivery.
Nie, L.
Show abstract
Compact tissue-specific promoters are highly desirable for gene therapy because viral vectors possess limited packaging capacity. However, existing promoter engineering strategies rely primarily on rational design or de novo sequence generation and lack efficient approaches for compressing long native promoters while preserving regulatory specificity. Although genome foundation models have substantially improved sequence-to-function prediction, they have not been effectively translated into computational platforms for promoter engineering. Here, we present VirEvo, a computational promoter engineering framework that integrates a virtual dual-luciferase assay (VirDLA), genome-foundation-model-guided genetic evolution, and an orthogonal Pan-Tissue Consistency Filter (PTCF). VirDLA introduces an internal-reference normalization strategy inspired by dual-luciferase reporter assays, enabling relative comparison of promoter activity across tissues without retraining AlphaGenome. Guided by these normalized activity scores, VirEvo iteratively optimizes promoter selectivity, off-target activity, and sequence length. Using the human p16INK4a promoter as a proof of concept, VirEvo evolved a compact synthetic promoter, SRP2M, of only 398 bp, representing an 85.9% reduction in sequence length. Experimental validation using dual-luciferase reporter assays in senescent IMR90 fibroblasts demonstrated that SRP2M retained 77% of wild-type senescence selectivity while reducing basal leakage to 52% of the wild-type level. Together, these results demonstrate the feasibility of genome-foundation-model-guided promoter engineering. VirEvo provides a generalizable framework for designing compact tissue-specific regulatory elements and extends the application of genome foundation models from functional prediction to synthetic regulatory engineering.
Irving, O. J.; Khan, C. J.; Albrecht, T.
Show abstract
DNA assembly is a cornerstone of synthetic biology, enabling the construction of bespoke genetic systems for applications ranging from metabolic engineering to DNA nanotechnology. Conventional Gibson Assembly (GA), the most widely used method, relies on 5' exonucleolytic resection and elevated temperatures ([~]50 {degrees}C), which together prevent the retention of 5' modifications and restrict compatibility with temperature-sensitive functionalities. Here, we report a DNA assembly strategy, 3 exonuclease-mediated low-temperature DNA assembly (3LTDA), which generates complementary 5' overhangs while preserving 5' end integrity. This approach enables the efficient assembly of blunt-ended, 5'-functionalised DNA fragments into both linear and circular constructs at ambient temperature (21 {degrees}C), with some assembly observed at temperatures as low as 4{degrees}C. We systematically optimise reaction conditions and demonstrate that this method supports efficient plasmid re-circularisation and multi-fragment assembly, including the construction of a [~]12.5 kbp plasmid from multiple DNA components. Comparative analysis across several DNA substrates shows that, under their respective optimal conditions, this approach matches or exceeds GA performance, improving assembly efficiency by up to 12.8%. Sequence analysis confirms high fidelity with no detectable base-pairing errors across assembled junctions. Crucially, this method preserves chemically functionalised 5' termini, enabling downstream conjugation and biochemical functionality. Retention of azide and biotin modifications was verified through fluorescence imaging, bead-based co-localisation, and enzymatic activity in ELISA-based assays. This is in contrast to GA-assembled controls, which showed complete loss of functionality under comparable conditions. We further assembled 5 kbp dsDNA using 3LTDA from four independent segments, three with different fluorescence reporters, and the fourth containing a biotin group for microparticle conjugation, each on the 5 end. Under fluorescence illumination, bead-bound DNA with all three fluorescence markers were detected. Conventional GA assembled constructs, on the other hand, failed to retain the reporter groups and the fluorescent images did not show the presence of any fluorescent markers. In addition to enhanced performance, the method could also reduce reagent cost and eliminate the need for elevated temperatures, simplifying workflows and expanding the applicability of multi-functionalised DNA constructs. Collectively, this work establishes 3LTDA as a robust, low-temperature alternative to conventional GA, with advantages for applications requiring precise chemical modification, temperature-sensitive components, or deployment outside conventional laboratory environments.
Xia, B.; Kalogriopoulos, N. A.; Wen, R.; Lane, Z. M.; Li, H.; Buitrago, N.; Lee, S.; Gao, R. D.; Ive, I.; Kim, Y.; Ting, A. Y.; Szablowski, J. O.
Show abstract
Detection of molecules with cell-based sensors allows for conversion of binding events into gene expression outputs. Here, we present a cell-based sensor that can detect extracellular double-stranded DNA. This sensor is based on an engineered receptor which we call Luminescent Ultrasensitive Nucleic Acid Reporter, or LUNAR. LUNAR is based on a recently developed Programmable Antigen-gated G-protein-coupled Engineered Receptor (PAGER). PAGERs are a genetic fusion of an auto-inhibitory peptide, a protein-binding domain, and a modified kappa opioid receptor. PAGERs are gated by two binding events. First, a protein ligand displaces an intramolecular inhibitor, Arodyn, then a second ligand activates the receptor. By replacing the protein-binding domain with a DNA binding zinc finger protein (ZFP) we could detect extracellular DNA in a dose-dependent fashion. Here, we show that first-generation LUNAR constructs can detect both oligonucleotides and plasmid double-stranded DNA with nanomolar sensitivity in mammalian cells. Future work will focus on improving sensitivity, fold-change, and multiplexing capabilities for sequence-specific DNA detection.
Jaiswal, B.; Black, T.; Namboothiri, H. R.; Pochana, K.; Hu, C. Y.
Show abstract
Optogenetic control enables light-actuated regulation of gene expression and provides a programmable interface between living cells and electronic systems. However, routine prototyping of optogenetic constructs remains limited by infrastructure. Existing closed-loop platforms often require chemostats, microfluidics, robotic handling, or custom optical sensors, which can increase cost, reduce accessibility, or constrain measurement performance. Here, we present LEMOS 2.0, an updated LED-Embedded Microplate for Optogenetic Studies, a low-cost device for optogenetic stimulation and gene-circuit characterization inside standard off-the-shelf microplate readers. LEMOS 2.0 builds on the original LEMOS platform by increasing throughput from 16 to 32 microwells and reducing light leakage between adjacent microwells, allowing dark conditions to be used as an additional illumination state. The device consists of a 3D-printed frame, individually addressable LEDs positioned next to each microwell, a rechargeable battery, and an onboard microcontroller for Bluetooth-based wireless communication. Biocompatible polydimethylsiloxane microwells are cast directly into the device by replica molding, allowing bacterial cultures to be stimulated while optical density and fluorescence are measured by the microplate reader. This protocol describes the full LEMOS 2.0 workflow, including device fabrication, circuit assembly, Arduino programming, PDMS microwell casting, plate-reader setup, strain and culture preparation, automated experiment execution, device cleanup, and fluorescence/OD600 data analysis. As a demonstration, the protocol uses the CcaSR optogenetic system, in which sfGFP expression is activated by green light and repressed by red light. LEMOS 2.0 is intended to make optogenetic perturbation and gene-expression characterization more accessible to wet-lab users, enabling faster design-build-test-learn cycles without requiring specialized bioreactor or microfluidic infrastructure.
Palmer, P.; Teran, N.; Wheeler, N.; Yassif, J. M.
Show abstract
As biological AI models become more powerful, practical biosecurity approaches are needed to support beneficial applications while reducing misuse risks. Sequence-similarity-based screening approaches are no longer adequate to safeguard biological AI models because these models can design molecules with novel sequences and structures. Therefore, a screening approach that takes function into account is needed. To address this need, we propose a new screening method for AI-enabled protein binder design tools. Our framework screens protein binding targets, with a focus on the human proteome, as opposed to the binder molecule itself. We constructed a database of 14,541 potentially harmful proteoform targets from the human proteome (7.1% of all human protein proteoforms) classified by biosecurity risk level. To discern structural and functional features, we evaluated constructs with an embedding-based screening method using the ESM-C protein language model. ESM-C achieved high accuracy for detecting variants of known targets (F1 scores >97%), with performance similar to BLASTP. However, ESM-C proved to be more effective at capturing functional relationships, distinguishing benign mutations from damaging ones where BLASTP did not. To characterize how screening would affect bioscience research, we measured flagging rates across diverse protein datasets. Flagging rates were significant for mammalian proteins weighted by publication frequency (23% for human, 20% for mouse), and rates for organisms distantly related to humans were minimal (<1.1% for bacteria, fungi, plants, and viruses). Among commercially relevant targets, 63% of antibody patent targets were classified as dual-use, reflecting that therapeutically important proteins often perform critical biological functions. To identify and flag risky user requests from protein binder design tools without placing an undue burden on scientific research and innovation, it will be essential to deploy this screening approach in a way that addresses the overlap our analysis showed between targets of concern and therapeutic targets-possibly in concert with tiered trusted access frameworks. This new method provides a foundation for proportionate safeguards for biological AI models that reduce misuse risks while preserving their benefits for legitimate research and demonstrates a concrete proof of principle that can be generalized to other protein design tools and biological AI models.
Ahavi, P.; Hoang, T.-N.-A.; Meyer, P.; Epaulard, O.; Le Gouellec, A.; Faulon, J.-L.
Show abstract
Although metabolomics has shown considerable promise for biomarker discovery, and the development of diagnostic and prognostic applications, its translation into routine clinical practice remains limited by analytical complexity, cost, throughput, and standardization challenges. These limitations underscore the need for complementary tools, particularly in resource-limited settings. In this study, we developed a workflow for the engineering and characterization of growth-coupled metabolic sensors capable of disease detection (healthy vs. infected) and outcome prediction (mild vs. severe), which we illustrated using COVID-19 as a proof-of-concept application. We first generated a biomarker-guided library of 34 candidate sensors leveraging both auxotrophic phenotypes and less stringent metabolic dependencies. We then screened the library against patient plasma pools, identifying 19 sensor candidates with diagnostic and/or prognostic potential, including 14 with prognostic potential. Lastly, a selected subset of candidates was further evaluated on a patient cohort using two newly developed analytical frameworks designed to extract additional information from bacterial growth curves. The best-performing sensors achieved a balanced accuracy of 0.88{+/-} 0.06 for prognostic prediction (outer-test AUC = 0.89, 5-fold cross-validation, n = 37) and 1.00 for diagnostic classification (outer-test AUC = 1.00, 5-fold cross-validation, n = 56). Collectively, these findings establish a proof of concept for translating disease-associated plasmatic metabolic signatures into low-cost, growth-coupled biosensors with diagnostic and prognostic capabilities.
Meeson, K.;Gaffney, R.;Schwartz, J.;Rattray, M.
Show abstract
There are huge variations in metabolic complexity between the different kingdoms of life. Whilst it has been shown that some simple, unicellular organisms such as E. coli direct their energetic resources towards maximising proliferation, the metabolic goals of more complex organisms are unclear. This is an especially important topic for engineered organisms, such as Chinese Hamster Ovary (CHO) cells, that have been modified to produce therapeutically relevant compounds. This metabolic goal is reflected in the objective function of a constraint-based model (CBM) and has a direct impact on the metabolic flux distribution that is predicted using Flux Balance Analysis (FBA). However, there is no broadly applicable approach to infer this objective function from experimental data, to ensure CBMs represent real growth conditions. Here, we developed SIMOFF (SIMulated annealing Objective Function Finder) to infer the most appropriate objective function from minimal experimental flux data. Our applications of SIMOFF to S. cerevisiae demonstrated that the most suitable objective function is dependent on key metabolic phenotypes, even when the same organism and conditions are being modelled. Furthermore, we demonstrated the translatability of SIMOFF through application to CHO cells, where we showed that a SIMOFF-inferred objective function improved the accuracy of gene essentiality simulations, resulting in more reliable experimental target predictions.
Hughes, N. W.; Kulkarni, S.; Goldman, G.; Marsiglia, J.; Jain, S.; Spees, K.; Hua Fu, B. X.; Vaalavirta, K.; Nakamura, M.
Show abstract
The problem of how protein sequences translate into defined functions remains largely unsolved despite decades of progress. New methods to efficiently explore protein sequence space will help to shed light on these sequence-function relationships, particularly for complex protein function. Here, we describe an approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM). We demonstrate the utility of this approach with the generation of novel compact RNA-guided nucleases. This approach is highly efficient, resulting in active nucleases with [~]40% sequence divergence relative to natural proteins and activity equivalent to or exceeding by up to [~]3X that of other compact nucleases at multiple endogenous loci in human cells. The approach described here is rapidly deployable and produces new sequences that will serve as scaffolds for further exploration of complex protein functionality, as well as substrates for novel genome engineering applications.